Papers with NLP techniques

35 papers
Assessing Post Deletion in Sina Weibo: Multi-modal Classification of Hot Topics (D19-50)

Copied to clipboard

Challenge: Weibo monitors and deletes posts to conform to government requirements . a recent study found that sentiment is the only indicator of censorship that is consistent across topics .
Approach: They analyze a dataset of censored and uncensore censors on Weibo . they use deep learning, CNN localization, and NLP techniques to analyze the data .
Outcome: The proposed analysis of censored and uncensoreded posts in Weibo shows that sentiment is the only indicator of a topic's censorship . censors can remove posts that are considered sensitive in three hours on average .
Integrating Ethics into the NLP Curriculum (2020.acl-tutorials)

Copied to clipboard

Challenge: a tutorial aims to teach students how to ethically apply NLP . the tutorial will focus on examples for university classrooms, but ideas can extend to company-internal workshops or tutorials in a variety of organizations.
Approach: a tutorial aims to empower NLP researchers and practitioners to teach others about ethical NLP . the tutorial will focus on examples for university classrooms, but ideas can extend to company-internal workshops .
Outcome: a tutorial aims to teach students how to ethically apply NLP techniques . the tutorial will include examples for university classrooms, but ideas can extend to company-internal workshops .
Designing, Evaluating, and Learning from Humans Interacting with NLP Models (2023.emnlp-tutorial)

Copied to clipboard

Challenge: This tutorial will cover how to conduct human-in-the-loop usability evaluations to ensure that models are capable of interacting with humans.
Approach: They will provide a systematic overview of key considerations and effective approaches for studying human-NLP model interactions.
Outcome: This tutorial will cover how to conduct human-in-the-loop usability evaluations to ensure that models are capable of interacting with humans.
Beyond Multiword Expressions: Processing Idioms and Metaphors (P18-5)

Copied to clipboard

Challenge: idioms and metaphors processing is a rapidly growing area in NLP, says dr. s. robertson . idiomatic idiomas are characteristic to all areas of human activity and to all types of discourse.
Approach: This tutorial will provide attendees with a clear notion of idioms and metaphors . it will provide them with computational models of linguistic characteristics and methods .
Outcome: This tutorial aims to provide attendees with a clear notion of the linguistic characteristics of idioms and metaphors . it outlines how to model idiomatic idiomes and their processing and what resources are available to support their use .
SEPSIS: I Can Catch Your Lies – A New Paradigm for Deception Detection (2025.acl-srw)

Copied to clipboard

Challenge: a new framework categorizes deception into three forms: lies of omission, lies of commission, and lies of influence . a novel framework for deception detection leveraging NLP techniques is proposed .
Approach: They propose a framework that categorizes deception into three forms: lies of omission, lies of commission, and lies of influence.
Outcome: The proposed framework achieves an impressive F1 score of 0.87 across all layers . it can be used to investigate lies of omission, lies of commission and lies of influence .
YATO: Yet Another deep learning based Text analysis Open toolkit (2023.emnlp-demo)

Copied to clipboard

Challenge: YATO is an open-source toolkit for text analysis with deep learning . it supports free combinations of three types of widely used features .
Approach: They introduce YATO, an open-source toolkit for text analysis with deep learning.
Outcome: YATO is an open-source toolkit for text analysis with deep learning . the toolkit supports free combinations of three types of widely used features .
Reaction Miner: An Integrated System for Chemical Reaction Extraction from Textual Data (2023.emnlp-demo)

Copied to clipboard

Challenge: Reaction Miner is a system designed to extract chemical reactions from raw scientific PDFs.
Approach: They propose a system that extracts chemical reactions directly from raw scientific PDFs.
Outcome: The proposed system can extract chemical reactions from raw scientific PDFs.
Towards Safer Operations: An Expert-involved Dataset of High-Pressure Gas Incidents for Preventing Future Failures (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing datasets for incident management tasks are labor-intensive and time-consuming.
Approach: They propose a new IncidentAI dataset for safety prevention that includes three tasks . they argue that NLP techniques are beneficial for analyzing incident reports .
Outcome: The proposed dataset shows that NLP techniques are beneficial for analyzing incident reports to prevent future failures.
The UIR Uncertainty Corpus for Chinese: Annotating Chinese Microblog Corpus for Uncertainty Identification from Social Media (L18-1)

Copied to clipboard

Challenge: Uncertainty identification is an important semantic processing task, critical to the quality of information in terms of factuality in many NLP techniques and applications.
Approach: They propose to annotate Chinese microblogs with an open uncertainty corpus . they propose to use contextual uncertain semantics rather than traditional cue-phrases to identify uncertainty .
Outcome: The proposed corpus can be used to identify uncertainty in social media texts.
Diffusion-NAT: Self-Prompting Discrete Diffusion for Non-Autoregressive Text Generation (2024.eacl-long)

Copied to clipboard

Challenge: Existing non-autoregressive (NAR) text-to-text generation methods are unable to generate coherent and fluent texts due to discrete nature of text.
Approach: They propose to integrate discrete diffusion models (DDM) into NAR text-to-text generation and integrate BART to improve the performance.
Outcome: The proposed method outperforms competing methods and surpasses autoregressive methods on 7 datasets.
Detecting Linguistic Characteristics of Alzheimer’s Dementia by Interpreting Neural Models (N18-2)

Copied to clipboard

Challenge: Current diagnoses often involve lengthy medical evaluations.
Approach: They apply neural models based on CNNs, LSTM-RNNs, and their combination to classify AD and control language samples.
Outcome: The proposed model achieves independent benchmark accuracy for the AD classification task.
A Survey on Natural Language Processing for Programming (2024.lrec-main)

Copied to clipboard

Challenge: Natural language processing for programming is a field of NLP and software engineering . it is used to assist programming, and is increasingly prevalent for its effectiveness in improving productivity.
Approach: They propose to use NLP techniques to assist programming by obtaining a structure-based representation and a functionality-oriented algorithm.
Outcome: The proposed approach could relieve developers from laborious work while improving efficiency for non-professional users.
DuoRC: Towards Complex Language Understanding with Paraphrased Reading Comprehension (P18-1)

Copied to clipboard

Challenge: DuoRC contains 186,089 unique question-answer pairs created from 7680 movie plots .
Approach: They propose a novel dataset for Reading Comprehension that motivates new challenges for neural approaches in language understanding beyond those offered by existing RC datasets.
Outcome: The proposed dataset motivates several new challenges for neural approaches in language understanding beyond those offered by existing RC datasets.
A Compliance Checking Framework Based on Retrieval Augmented Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing text-based compliance checking methods are limited by their flexibility and lack structure.
Approach: They propose a text-based compliance checking framework based on Retrieval-Augmented Generation that integrates a static layer for storing factual knowledge, a dynamic layer for retrieval and reasoning, and an eventic graph to structurally describe regulatory information.
Outcome: The proposed framework consistently achieves state-of-the-art results across various scenarios surpassing baselines.
Challenges and Opportunities of Applying Natural Language Processing in Business Process Management (C18-1)

Copied to clipboard

Challenge: In the last decade, the maturity achieved by NLP has turned the spotlight to the possibilities offered by Nlp for a variety of novel applications.
Approach: They propose to use NLP to raise the benefits of BPM at different levels . they propose to focus on the daily tasks that an organization must perform .
Outcome: The proposed approach could be applied to a variety of business processes in a scalable fashion.
The Promises and Pitfalls of Using Language Models to Measure Instruction Quality in Education (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to assess instruction quality require trained raters to observe classrooms based on established criteria.
Approach: They propose to use Natural Language Processing techniques to assess multiple high-inference instructional practices in in-person K-12 classrooms and simulated performance tasks for pre-service teachers.
Outcome: The proposed method is able to assess multiple high-inference instructional practices in two educational settings: in-person K-12 classrooms and simulated performance tasks for pre-service teachers.
Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where? (2025.emnlp-main)

Copied to clipboard

Challenge: 20% of all papers in the ACL Anthology address social good issues . authors are more likely to do work addressing social good concerns when publishing in venues outside of ACL.
Approach: They use author- and venue-level perspectives to map the landscape of NLP4SG . they find authors are more likely to do work addressing social good concerns outside of ACL .
Outcome: The study analyzes the literature on NLP4SG and its impact on the ACL community . 20% of all papers in the anthology address social good issues, the study finds .
Human-in-the-loop Robotic Grasping Using BERT Scene Representation (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches for robotic grasping in cluttered scenes are expensive and lack structure information.
Approach: They propose a human-in-the-loop framework for robotic grasping in cluttered scenes . they substitute scene-graph representation with a text representation of the scene using BERT .
Outcome: The proposed framework outperforms object-agnostic and scene-graph based methods on robots and physical robots.
Acquiring Social Knowledge about Personality and Driving-related Behavior (2020.lrec-1)

Copied to clipboard

Challenge: Using crowdsourcing, we acquire human-specific knowledge about personality and driving.
Approach: They propose a psychological approach to collect human-specific social knowledge from a text corpus using NLP techniques.
Outcome: The proposed approach collects human-specific social knowledge from a text corpus, and then implements it into a system.
Lexical and Semantic Features for Cross-lingual Text Reuse Classification: an Experiment in English and Latin Paraphrases (L18-1)

Copied to clipboard

Challenge: Analyzing historical languages is challenging because they lack primary material for certain time periods . under-resourced languages such as Ancient Greek and Latin lack advanced natural-language processing (NLP) techniques .
Approach: They propose to use machine learning to detect and classify paraphrastic text reuse in historical texts.
Outcome: The proposed method improves the accuracy of paraphrastic text reuse detection in historical languages.
Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model Prompting (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on the integration of Differential Privacy (DP) into NLP techniques.
Approach: They propose a method for text privatization leveraging language models to rewrite texts . they examine the usability of DP in NLP and its benefits over non-DP approaches .
Outcome: The proposed method is a novel method for text privatization leveraging language models to rewrite texts.
BanFakeNews: A Dataset for Detecting Fake News in Bangla (2020.lrec-1)

Copied to clipboard

Challenge: Impact of fake news is creating havoc worldwide.
Approach: They propose an annotated dataset of 50K news that can be used for building automated fake news detection systems for a low resource language like Bangla.
Outcome: The proposed system can be built with state-of-the-art NLP techniques for a low resource language like Bangla.
CrisiText: A dataset of warning messages for LLM training in emergency communication (2026.findings-eacl)

Copied to clipboard

Challenge: Identifying threats and mitigating their potential damage during crisis situations is paramount for safeguarding endangered individuals.
Approach: They present a large-scale dataset for the generation of warning messages across 13 different types of crisis scenarios.
Outcome: The proposed dataset contains more than 400,000 warning messages (spanning almost 18,000 crisis situations) aimed at assisting civilians during and after such events.
ESCRITO - An NLP-Enhanced Educational Scoring Toolkit (L18-1)

Copied to clipboard

Challenge: Existing implementations are very specific to specific use cases and datasets.
Approach: ESCRITO is a toolkit for scoring student writings using NLP techniques . authors propose teachers and NLP researchers to use APIs for scoring pipelines .
Outcome: ESCRITO is a toolkit for scoring student writings using NLP techniques . it addresses two main user groups: teachers and NLP researchers .
MuLD: The Multitask Long Document Benchmark (2022.lrec-1)

Copied to clipboard

Challenge: Existing benchmarks for NLP focus on tasks for one or two sentences, but efficient techniques are needed for processing much longer sequences.
Approach: They propose to modify existing NLP tasks to create a long document benchmark which requires models to successfully model long-term dependencies in the text.
Outcome: The proposed benchmark is much more challenging than its ‘short document’ equivalents.
Detection and Positive Reconstruction of Cognitive Distortion Sentences: Mandarin Dataset and Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have investigated the application of NLP models in English for each stage of this process.
Approach: They propose a Positive Reconstruction Framework based on broaden-and-build theory to address and reframe negative thoughts through a positive reinterpretation.
Outcome: The proposed framework is based on broaden-and-build theory and can detect cognitive distortions and suggest a positive reframe in Mandarin.
Exploiting Commonsense Knowledge about Objects for Visual Activity Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing tasks that aim to identify the objects in an image are object detection and image classification, but recent work has focused on more comprehensive image under- standing tasks.
Approach: They propose to incorporate commonsense knowledge about physical objects into a transformer-based model that is trained to predict the actionverb for visual activity recognition.
Outcome: The proposed model incorporates prototypical function knowledge about physical objects to predict the actionverb for visual activity recognition.
EmoMent: An Emotion Annotated Mental Health Corpus from Two South Asian Countries (2022.coling-1)

Copied to clipboard

Challenge: Recent research using AI and NLP demonstrates strong potential to automatically detect mental health issues from digital footprints such that professionals could provide timely interventions and mental health resources to vulnerable persons.
Approach: They developed an emotion-annotated mental health corpus from 2802 Facebook posts extracted from two South Asian countries, Sri Lanka and India.
Outcome: The proposed model achieved 98.3% agreement between the annotators and a Fleiss’ Kappa of 0.82.
Survey on Thai NLP Language Resources and Tools (2022.lrec-1)

Copied to clipboard

Challenge: Thai language is one of the under-resourced languages in the NLP domain, although it is spoken by approximately 70 million people globally.
Approach: They propose to use Thai language as an example to understand how NLP works and how it can be applied to Thai language.
Outcome: The results show that Thai NLP research has progressed over the past three decades, especially on upstream tasks such as tokenisation, but research on downstream tasks such syntactic parsing and semantic analysis is still limited.
A Fine-grained Chinese Software Privacy Policy Dataset for Sequence Labeling and Regulation Compliant Identification (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets that ignore law requirements are limited to English.
Approach: They construct a Chinese privacy policy dataset that can be used to analyze software privacy policies.
Outcome: The proposed dataset includes 483 Chinese Android privacy policies, over 11K sentences, and 52K fine-grained annotations.
Qualitative Code Suggestion: A Human-Centric Approach to Qualitative Coding (2023.findings-emnlp)

Copied to clipboard

Challenge: Qualitative coding is a content analysis method that assigns descriptive labels or qualitative codes to passages.
Approach: They propose a qualitative code suggestion task where a ranked list of previously assigned qualitative codes is suggested from an identified passage.
Outcome: The proposed method integrates previously ignored properties such as the sequence in which passages are annotated, the importance of rare codes and the differences in annotation styles between coders.
Don’t Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections (2025.acl-long)

Copied to clipboard

Challenge: Cultural Heritage metadata can contain outdated or offensive terms that reflect historical cultural and societal norms.
Approach: They propose an AI-powered tool that detects offensive terms in CH metadata . the tool has processed over 7.9 million records and provides contextual insights .
Outcome: The proposed tool has processed over 7.9 million records and provides contextual insights . it pairs biased language with contextual information and suggestions for appropriate usage .
Reconstruction of Cuneiform Literary Texts as Text Matching (2024.lrec-main)

Copied to clipboard

Challenge: cuneiform fragment identification is a slow and unsystematic process for reconstructing ancient texts . fragments of cuniform script are often found in fragments written on clay tablets . cnl is able to identify fragments and match them with existing text collections .
Approach: They propose a character-level n-gram-based similarity matching approach to identify fragments . they compare different approaches to identify overlaps between fragments and texts .
Outcome: The proposed approach speeds up the process and reduces the time it takes to complete the work.
Towards Comprehensive Language Analysis for Clinically Enriched Spontaneous Dialogue (2024.lrec-main)

Copied to clipboard

Challenge: Contemporary NLP has progressed from feature-based classification to fine-tuning and prompt-based techniques . many of these techniques remain understudied in the context of real-world, clinically enriched spontaneous dialogue.
Approach: They investigate the efficacy and overall performance of a range of NLP techniques on transcribed speech from patients with schizophrenia and other disorders.
Outcome: The proposed methods are effective in analyzing transcribed speech from patients with schizophrenia and healthy controls taking a clinically-validated language test.
Investigating Links between Illicit Massage Businesses through Natural Language Processing and Graph Machine Learning (2026.findings-acl)

Copied to clipboard

Challenge: Illicit massage businesses exploit vulnerable individuals through forced sex or labor . identifying key indicators from vast volume of data associated with these businesses poses significant challenge .
Approach: They propose a multi-stream data integration approach focusing on Yelp reviews . they propose bespoke subgraph extraction strategies to detect links between massage businesses .
Outcome: The proposed approach outperforms baseline methods in a multi-stream data integration framework based on consumer reviews on Yelp.com and contextual data from the U.S. Census and business license records.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations